Install coherent deployment plans with dataset-bound SDS identities - #749
Merged
Merged
Conversation
zzylol
marked this pull request as ready for review
September 21, 2026 16:13
zzylol
added a commit
that referenced
this pull request
Sep 22, 2026
Integrate PR #749 reader/writer bindings and maintenance projections while preserving data-partition worker ownership and plan-derived configuration. Keep selected and maintenance DAG schemas distinct and adapt shared-sink execution to the projected graph.
This was referenced Sep 22, 2026
zzylol
force-pushed
the
docs/physical-plan-design
branch
from
September 28, 2026 13:47
b0d77ce to
5f1eebf
Compare
zzylol
force-pushed
the
refactor/backend-plan-split
branch
from
September 28, 2026 16:14
12896bd to
70a8d88
Compare
CompileAndPublishPhysicalPlanRequest gained a required `dataset_identity` field in this branch, but two api_tests fixtures build their request body as JSON by hand and were never updated, so both failed deserialization with `missing field dataset_identity` before reaching the handler. Take the value from the planning snapshot's own `environment.dataset_identity` rather than inventing one, so the fixture keeps describing the same dataset the rest of the snapshot describes. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
zzylol
added a commit
that referenced
this pull request
Sep 28, 2026
…774) * docs: clarify Planner physical plan and SDS architecture * docs: specify executable subplan materialization boundaries * docs: scope migration to backend precompute and query plans * docs: clarify window terminology migration * Revert "docs: clarify window terminology migration" This reverts commit daa5281. * docs: focus physical plan and SDS designs * docs: add physical compiler input example * docs: add physical compiler output example * docs: include query expression in compiler example * docs: reorganize physical plan integration design * docs: reorganize SDS and migration designs * docs: define maintenance inputs before plan split example * docs: use plan version consistently in backend design * docs: remove standalone catalog materialization abstraction * docs: add concise planner backend glossary * refactor: split installed maintenance DAGs from query execution * refactor: remove backend Collector dependency and normalize legacy DAGs * test: restore whole-backend process coverage without Collector * fix: migrate Planner main and reject legacy runtime artifacts * test: send full sketch envelope in whole-backend E2E * refactor: bind backend summary state through versioned SDS slots * test: send complete sketch envelopes in process fixtures * docs: clarify selected deployment guarantee terminology * docs: explain missing planner maintenance guarantee * docs: motivate selected producer maintenance decision * docs: label catalog reads and SDS metadata ownership * docs: align SDS ownership and lifecycle terminology * docs: use current planner and plan-version names consistently * docs: align integration diagram with SDS ownership * docs: clarify instance identity and shared producer wording * refactor: align plan bindings with current Planner and SDS contract * docs: distinguish summary definitions from runtime stores * docs: model one runtime summary store for DAG bindings * docs: scope SDS lifecycle to read eligibility * docs: tie stored summary examples directly to DAG outputs * docs: name summary tables, stored records, and output references by role * docs: illustrate summary definitions, stored records, and output references * docs: limit v1 summary storage to definitions and stored summaries * refactor: align plan bindings with summary store v1 * fix: validate selected DAG provenance * fix: retain neutral sketch codec dependencies when syncing main * fix: align derived DAG validation with current schema versions * fix: keep maintenance document version distinct from complete DAG version * refactor: adopt costed Planner selection without legacy API adapters * test: verify exact process routing for uncertified Planner candidates * docs: remove redundant Planner selection contract * deps: pin Planner bounded HLL confidence model * test: size transmitted KLL state from a certified accuracy contract * test: retain KLL collector capability when using theoretical confidence * refactor: implement typed summary semantics and physical plan lowering * docs: separate Planner physical computation from backend deployment * refactor: name the backend orchestration entry DeploymentPlanCompiler * Restack PR #770 with implementation before standalone acceptance * chore: consume Planner physical precompute candidate interfaces * chore: consume Planner materialization frontier enumeration * chore: consume exact temporal ranking physical candidates * chore: use shared Planner candidate winner selection * chore: consume sparse counter shared readout contracts * test: declare collector ranking fixture capabilities explicitly * Use Planner counter-window candidate execution * Use shared keyed-counter omission contract * Use Planner exact-counter population omission * Construct fixture evidence for the standalone operator foundation * build: pin Planner exact-state scratch merge implementation * build: pin shared finalized-pane reconstruction fix * docs: bind Planner physical DAGs without backend re-lowering * docs: describe summary inputs with groups and pane duration * docs: define summary semantic completeness beyond input scope * docs: define SDS identity through canonical Planner computation * docs: decouple SDS semantic identity from executable Planner IR * docs: define Planner-owned SDS discovery for future ad hoc queries * docs: streamline SDS design around definitions and stored results * docs: track bound-query SDS migration across implementation PRs * docs: identify active shared-library PR in bound-query migration * build: align shared Planner dependencies with remote PR 462 * docs: separate bound SDS range lookup from state validation * feat(sds): separate bound output routing from persisted semantic identity * docs: state SDS migration responsibilities without stale implementation claims * refactor(sds): keep semantic variants compact without changing wire format * test: align evidence fixtures with selected exact count state * docs: align precompute SDS description with semantic catalog * test: bind imported-state fixtures to their actual semantic definition * test: distinguish imported CMS transport from total-count planning * docs: clarify Planner physical plan and SDS architecture * docs: specify executable subplan materialization boundaries * docs: scope migration to backend precompute and query plans * docs: clarify window terminology migration * Revert "docs: clarify window terminology migration" This reverts commit daa5281. * docs: focus physical plan and SDS designs * docs: add physical compiler input example * docs: add physical compiler output example * docs: include query expression in compiler example * docs: reorganize physical plan integration design * docs: reorganize SDS and migration designs * docs: define maintenance inputs before plan split example * docs: use plan version consistently in backend design * docs: remove standalone catalog materialization abstraction * docs: add concise planner backend glossary * docs: clarify selected deployment guarantee terminology * docs: explain missing planner maintenance guarantee * docs: motivate selected producer maintenance decision * docs: label catalog reads and SDS metadata ownership * docs: align SDS ownership and lifecycle terminology * docs: use current planner and plan-version names consistently * docs: align integration diagram with SDS ownership * docs: clarify instance identity and shared producer wording * docs: distinguish summary definitions from runtime stores * docs: model one runtime summary store for DAG bindings * docs: scope SDS lifecycle to read eligibility * docs: tie stored summary examples directly to DAG outputs * docs: name summary tables, stored records, and output references by role * docs: illustrate summary definitions, stored records, and output references * docs: limit v1 summary storage to definitions and stored summaries * docs: separate Planner physical computation from backend deployment * docs: bind Planner physical DAGs without backend re-lowering * docs: describe summary inputs with groups and pane duration * docs: define summary semantic completeness beyond input scope * docs: define SDS identity through canonical Planner computation * docs: decouple SDS semantic identity from executable Planner IR * docs: define Planner-owned SDS discovery for future ad hoc queries * docs: streamline SDS design around definitions and stored results * docs: track bound-query SDS migration across implementation PRs * docs: identify active shared-library PR in bound-query migration * docs: separate bound SDS range lookup from state validation * docs: state SDS migration responsibilities without stale implementation claims * docs: preserve design index after rebasing onto main * docs: clarify SDS source identity and version-scoped recovery * docs: illustrate SDS identity and recovery decisions * feat: bind Planner dataset semantics to deployment input identity * fix: recognize shared native batch encoding at dependency boundary * test: reject pre-dataset planning snapshot versions * refactor: version dataset-bound catalog and update empty-plan fixtures * docs: describe dataset-bound planning and installation inputs * docs: clarify candidate selection and deployment ownership * refactor: split backend deployment plans and bind stored outputs * fix: restrict foundation SDS recovery to the installed generation * fix: allocate fresh physical series when a plan version changes * test: record foundation rebase and recovery regression evidence * fix: complete shared state encoding adoption at the dependency boundary * style: satisfy workspace formatting after the compiler rename * feat(sds): separate bound output routing from persisted semantic identity * feat(sds): require dataset-bound Planner definitions and versioned catalogs * refactor: implement typed summary semantics and physical plan lowering * fix: validate final SDS semantics across Planner adapters and recovery * test: verify final SDS identity installation, serving and restart * test: align shared execution fixtures with final #749 SDS foundation * docs: record shared execution directly after the SDS foundation * fix(control-plane): supply dataset_identity in the API test fixtures CompileAndPublishPhysicalPlanRequest gained a required `dataset_identity` field in this branch, but two api_tests fixtures build their request body as JSON by hand and were never updated, so both failed deserialization with `missing field dataset_identity` before reaching the handler. Take the value from the planning snapshot's own `environment.dataset_identity` rather than inventing one, so the fixture keeps describing the same dataset the rest of the snapshot describes. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * Remove obsolete SDS wire aliases and persisted metadata migration * fix(data-plane): initialize dataset_identity in the bootstrap ingest contract IngestContract gained an optional `dataset_identity`, but the bootstrap envelope in data_plane's entry point was not updated, so the binary failed to compile with E0063 and took the whole workspace build with it. The bootstrap envelope describes the state before any plan is installed, where every other field is a placeholder, so there is no dataset to bind to yet; `None` is the accurate value. A published plan supplies the identity. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * Reject obsolete state-column syntax in compiler roundtrip coverage * Remove unused materialization wire adapter and propagate recovery errors * Keep projection fixtures on the canonical typed encoding * Remove superseded deployment API field aliases * fix(control-plane): supply dataset_identity in the API test fixtures CompileAndPublishPhysicalPlanRequest gained a required `dataset_identity` field in this branch, but two api_tests fixtures build their request body as JSON by hand and were never updated, so both failed deserialization with `missing field dataset_identity` before reaching the handler. Take the value from the planning snapshot's own `environment.dataset_identity` rather than inventing one, so the fixture keeps describing the same dataset the rest of the snapshot describes. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> (cherry picked from commit e627dbb) * Document obsolete-format removals and cleanup validation * Record passing restacked plan, storage and serving checks * Remove generated evaluation archives and keep validation summaries * Normalize validation summary formatting * Remove PR process reports and defer execution design to shared-runtime PR * Keep discovery and calibration on dataset-bound snapshot version 3 * Use typed deployment configuration in process E2E fixtures * Align process assertions with dataset-bound SDS and cold successor activation * fix: execute typed Planner fragments and preserve terminal resource errors * fix: bind exact integer samples without losing input types * test: verify exact integer protocol input binding * refactor: name per-boundary physical fragments explicitly * fix: consume finalized Planner query candidate outputs * test: require Planner filters for bound protocol vectors * test: use Planner schema lifting module * fix: bind Planner filters and finalized query results * docs: describe bound Planner filter execution * fix: retain state binding and sharing beneath explicit query readouts * fix: pin Planner query finalization for every candidate entry point * docs: distinguish query values from stored accumulator boundaries * test: assert exact Count state beneath its query readout --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
zzylol
added a commit
that referenced
this pull request
Sep 29, 2026
…eads (#763) * docs: remove redundant Planner selection contract * deps: pin Planner bounded HLL confidence model * test: size transmitted KLL state from a certified accuracy contract * test: retain KLL collector capability when using theoretical confidence * refactor: implement typed summary semantics and physical plan lowering * docs: separate Planner physical computation from backend deployment * refactor: name the backend orchestration entry DeploymentPlanCompiler * Restack PR #770 with implementation before standalone acceptance * Restack PR #763 with implementation before standalone acceptance * chore: consume Planner physical precompute candidate interfaces * chore: consume Planner materialization frontier enumeration * chore: consume exact temporal ranking physical candidates * chore: use shared Planner candidate winner selection * chore: consume sparse counter shared readout contracts * test: declare collector ranking fixture capabilities explicitly * Use Planner counter-window candidate execution * Use shared keyed-counter omission contract * Use Planner exact-counter population omission * Construct fixture evidence for the standalone operator foundation * build: pin Planner exact-state scratch merge implementation * build: pin shared finalized-pane reconstruction fix * docs: bind Planner physical DAGs without backend re-lowering * docs: describe summary inputs with groups and pane duration * docs: define summary semantic completeness beyond input scope * docs: define SDS identity through canonical Planner computation * docs: decouple SDS semantic identity from executable Planner IR * docs: define Planner-owned SDS discovery for future ad hoc queries * docs: streamline SDS design around definitions and stored results * docs: track bound-query SDS migration across implementation PRs * docs: identify active shared-library PR in bound-query migration * build: align shared Planner dependencies with remote PR 462 * docs: separate bound SDS range lookup from state validation * feat(sds): separate bound output routing from persisted semantic identity * docs: state SDS migration responsibilities without stale implementation claims * refactor(sds): keep semantic variants compact without changing wire format * test: align evidence fixtures with selected exact count state * docs: align precompute SDS description with semantic catalog * test: bind imported-state fixtures to their actual semantic definition * test: bind imported-state fixtures to their actual semantic definition * test: distinguish imported CMS transport from total-count planning * docs: clarify Planner physical plan and SDS architecture * docs: specify executable subplan materialization boundaries * docs: scope migration to backend precompute and query plans * docs: clarify window terminology migration * Revert "docs: clarify window terminology migration" This reverts commit daa5281. * docs: focus physical plan and SDS designs * docs: add physical compiler input example * docs: add physical compiler output example * docs: include query expression in compiler example * docs: reorganize physical plan integration design * docs: reorganize SDS and migration designs * docs: define maintenance inputs before plan split example * docs: use plan version consistently in backend design * docs: remove standalone catalog materialization abstraction * docs: add concise planner backend glossary * docs: clarify selected deployment guarantee terminology * docs: explain missing planner maintenance guarantee * docs: motivate selected producer maintenance decision * docs: label catalog reads and SDS metadata ownership * docs: align SDS ownership and lifecycle terminology * docs: use current planner and plan-version names consistently * docs: align integration diagram with SDS ownership * docs: clarify instance identity and shared producer wording * docs: distinguish summary definitions from runtime stores * docs: model one runtime summary store for DAG bindings * docs: scope SDS lifecycle to read eligibility * docs: tie stored summary examples directly to DAG outputs * docs: name summary tables, stored records, and output references by role * docs: illustrate summary definitions, stored records, and output references * docs: limit v1 summary storage to definitions and stored summaries * docs: separate Planner physical computation from backend deployment * docs: bind Planner physical DAGs without backend re-lowering * docs: describe summary inputs with groups and pane duration * docs: define summary semantic completeness beyond input scope * docs: define SDS identity through canonical Planner computation * docs: decouple SDS semantic identity from executable Planner IR * docs: define Planner-owned SDS discovery for future ad hoc queries * docs: streamline SDS design around definitions and stored results * docs: track bound-query SDS migration across implementation PRs * docs: identify active shared-library PR in bound-query migration * docs: separate bound SDS range lookup from state validation * docs: state SDS migration responsibilities without stale implementation claims * docs: preserve design index after rebasing onto main * docs: clarify SDS source identity and version-scoped recovery * docs: illustrate SDS identity and recovery decisions * feat: bind Planner dataset semantics to deployment input identity * fix: reject cross-version stored-state adoption * fix: recognize shared native batch encoding at dependency boundary * test: reject pre-dataset planning snapshot versions * refactor: version dataset-bound catalog and update empty-plan fixtures * docs: describe dataset-bound planning and installation inputs * docs: clarify candidate selection and deployment ownership * refactor: split backend deployment plans and bind stored outputs * fix: restrict foundation SDS recovery to the installed generation * fix: allocate fresh physical series when a plan version changes * fix: reserve the persisted storage handle during recovery * test: record foundation rebase and recovery regression evidence * fix: complete shared state encoding adoption at the dependency boundary * style: satisfy workspace formatting after the compiler rename * feat(sds): separate bound output routing from persisted semantic identity * feat(sds): require dataset-bound Planner definitions and versioned catalogs * refactor: implement typed summary semantics and physical plan lowering * fix: validate final SDS semantics across Planner adapters and recovery * test: verify final SDS identity installation, serving and restart * test: align shared execution fixtures with final #749 SDS foundation * docs: record shared execution directly after the SDS foundation * fix(control-plane): supply dataset_identity in the API test fixtures CompileAndPublishPhysicalPlanRequest gained a required `dataset_identity` field in this branch, but two api_tests fixtures build their request body as JSON by hand and were never updated, so both failed deserialization with `missing field dataset_identity` before reaching the handler. Take the value from the planning snapshot's own `environment.dataset_identity` rather than inventing one, so the fixture keeps describing the same dataset the rest of the snapshot describes. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * Remove obsolete SDS wire aliases and persisted metadata migration * fix(data-plane): initialize dataset_identity in the bootstrap ingest contract IngestContract gained an optional `dataset_identity`, but the bootstrap envelope in data_plane's entry point was not updated, so the binary failed to compile with E0063 and took the whole workspace build with it. The bootstrap envelope describes the state before any plan is installed, where every other field is a placeholder, so there is no dataset to bind to yet; `None` is the accurate value. A published plan supplies the identity. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * Reject obsolete state-column syntax in compiler roundtrip coverage * Remove unused materialization wire adapter and propagate recovery errors * Keep projection fixtures on the canonical typed encoding * Remove superseded deployment API field aliases * fix(control-plane): supply dataset_identity in the API test fixtures CompileAndPublishPhysicalPlanRequest gained a required `dataset_identity` field in this branch, but two api_tests fixtures build their request body as JSON by hand and were never updated, so both failed deserialization with `missing field dataset_identity` before reaching the handler. Take the value from the planning snapshot's own `environment.dataset_identity` rather than inventing one, so the fixture keeps describing the same dataset the rest of the snapshot describes. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> (cherry picked from commit e627dbb) * fix(control-plane): supply dataset_identity in the API test fixtures CompileAndPublishPhysicalPlanRequest gained a required `dataset_identity` field in this branch, but two api_tests fixtures build their request body as JSON by hand and were never updated, so both failed deserialization with `missing field dataset_identity` before reaching the handler. Take the value from the planning snapshot's own `environment.dataset_identity` rather than inventing one, so the fixture keeps describing the same dataset the rest of the snapshot describes. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> (cherry picked from commit e627dbb) * Document obsolete-format removals and cleanup validation * Record passing restacked plan, storage and serving checks * Remove generated evaluation archives and keep validation summaries * Normalize validation summary formatting * Remove PR process reports and defer execution design to shared-runtime PR * Keep discovery and calibration on dataset-bound snapshot version 3 * Use typed deployment configuration in process E2E fixtures * Remove unused streaming-config fixture after physical-only startup * Align process assertions with dataset-bound SDS and cold successor activation * fix: execute typed Planner fragments and preserve terminal resource errors * fix: bind exact integer samples without losing input types * test: verify exact integer protocol input binding * refactor: name per-boundary physical fragments explicitly * fix: consume finalized Planner query candidate outputs * test: require Planner filters for bound protocol vectors * test: use Planner schema lifting module * fix: bind Planner filters and finalized query results * docs: describe bound Planner filter execution * fix: retain state binding and sharing beneath explicit query readouts * fix: pin Planner query finalization for every candidate entry point * docs: distinguish query values from stored accumulator boundaries * test: assert exact Count state beneath its query readout * fix: preserve maintenance failures and stop failed worker admission * docs: specify bounded revisions and consistent query snapshots * feat: execute bounded precompute revisions with consistent SDS reads --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
zzylol
added a commit
that referenced
this pull request
Sep 29, 2026
…765) * refactor(sds): keep semantic variants compact without changing wire format * test: align evidence fixtures with selected exact count state * docs: align precompute SDS description with semantic catalog * test: bind imported-state fixtures to their actual semantic definition * test: bind imported-state fixtures to their actual semantic definition * test: distinguish imported CMS transport from total-count planning * docs: clarify Planner physical plan and SDS architecture * docs: specify executable subplan materialization boundaries * docs: scope migration to backend precompute and query plans * docs: clarify window terminology migration * Revert "docs: clarify window terminology migration" This reverts commit daa5281. * docs: focus physical plan and SDS designs * docs: add physical compiler input example * docs: add physical compiler output example * docs: include query expression in compiler example * docs: reorganize physical plan integration design * docs: reorganize SDS and migration designs * docs: define maintenance inputs before plan split example * docs: use plan version consistently in backend design * docs: remove standalone catalog materialization abstraction * docs: add concise planner backend glossary * docs: clarify selected deployment guarantee terminology * docs: explain missing planner maintenance guarantee * docs: motivate selected producer maintenance decision * docs: label catalog reads and SDS metadata ownership * docs: align SDS ownership and lifecycle terminology * docs: use current planner and plan-version names consistently * docs: align integration diagram with SDS ownership * docs: clarify instance identity and shared producer wording * docs: distinguish summary definitions from runtime stores * docs: model one runtime summary store for DAG bindings * docs: scope SDS lifecycle to read eligibility * docs: tie stored summary examples directly to DAG outputs * docs: name summary tables, stored records, and output references by role * docs: illustrate summary definitions, stored records, and output references * docs: limit v1 summary storage to definitions and stored summaries * docs: separate Planner physical computation from backend deployment * docs: bind Planner physical DAGs without backend re-lowering * docs: describe summary inputs with groups and pane duration * docs: define summary semantic completeness beyond input scope * docs: define SDS identity through canonical Planner computation * docs: decouple SDS semantic identity from executable Planner IR * docs: define Planner-owned SDS discovery for future ad hoc queries * docs: streamline SDS design around definitions and stored results * docs: track bound-query SDS migration across implementation PRs * docs: identify active shared-library PR in bound-query migration * docs: separate bound SDS range lookup from state validation * docs: state SDS migration responsibilities without stale implementation claims * docs: preserve design index after rebasing onto main * docs: clarify SDS source identity and version-scoped recovery * docs: illustrate SDS identity and recovery decisions * feat: bind Planner dataset semantics to deployment input identity * fix: reject cross-version stored-state adoption * fix: recognize shared native batch encoding at dependency boundary * test: reject pre-dataset planning snapshot versions * refactor: version dataset-bound catalog and update empty-plan fixtures * test: enforce whole-query consistency and new-version warm-up * docs: describe dataset-bound planning and installation inputs * docs: clarify candidate selection and deployment ownership * docs: align installation guide with dataset and recovery contracts * refactor: split backend deployment plans and bind stored outputs * fix: restrict foundation SDS recovery to the installed generation * fix: allocate fresh physical series when a plan version changes * fix: reserve the persisted storage handle during recovery * test: record foundation rebase and recovery regression evidence * fix: complete shared state encoding adoption at the dependency boundary * style: satisfy workspace formatting after the compiler rename * feat(sds): separate bound output routing from persisted semantic identity * feat(sds): require dataset-bound Planner definitions and versioned catalogs * refactor: implement typed summary semantics and physical plan lowering * fix: validate final SDS semantics across Planner adapters and recovery * test: verify final SDS identity installation, serving and restart * test: align shared execution fixtures with final #749 SDS foundation * docs: record shared execution directly after the SDS foundation * fix(control-plane): supply dataset_identity in the API test fixtures CompileAndPublishPhysicalPlanRequest gained a required `dataset_identity` field in this branch, but two api_tests fixtures build their request body as JSON by hand and were never updated, so both failed deserialization with `missing field dataset_identity` before reaching the handler. Take the value from the planning snapshot's own `environment.dataset_identity` rather than inventing one, so the fixture keeps describing the same dataset the rest of the snapshot describes. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * Remove obsolete SDS wire aliases and persisted metadata migration * fix(data-plane): initialize dataset_identity in the bootstrap ingest contract IngestContract gained an optional `dataset_identity`, but the bootstrap envelope in data_plane's entry point was not updated, so the binary failed to compile with E0063 and took the whole workspace build with it. The bootstrap envelope describes the state before any plan is installed, where every other field is a placeholder, so there is no dataset to bind to yet; `None` is the accurate value. A published plan supplies the identity. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * Reject obsolete state-column syntax in compiler roundtrip coverage * Remove unused materialization wire adapter and propagate recovery errors * Keep projection fixtures on the canonical typed encoding * Remove superseded deployment API field aliases * fix(control-plane): supply dataset_identity in the API test fixtures CompileAndPublishPhysicalPlanRequest gained a required `dataset_identity` field in this branch, but two api_tests fixtures build their request body as JSON by hand and were never updated, so both failed deserialization with `missing field dataset_identity` before reaching the handler. Take the value from the planning snapshot's own `environment.dataset_identity` rather than inventing one, so the fixture keeps describing the same dataset the rest of the snapshot describes. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> (cherry picked from commit e627dbb) * fix(control-plane): supply dataset_identity in the API test fixtures CompileAndPublishPhysicalPlanRequest gained a required `dataset_identity` field in this branch, but two api_tests fixtures build their request body as JSON by hand and were never updated, so both failed deserialization with `missing field dataset_identity` before reaching the handler. Take the value from the planning snapshot's own `environment.dataset_identity` rather than inventing one, so the fixture keeps describing the same dataset the rest of the snapshot describes. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> (cherry picked from commit e627dbb) * fix(control-plane): supply dataset_identity in the API test fixtures CompileAndPublishPhysicalPlanRequest gained a required `dataset_identity` field in this branch, but two api_tests fixtures build their request body as JSON by hand and were never updated, so both failed deserialization with `missing field dataset_identity` before reaching the handler. Take the value from the planning snapshot's own `environment.dataset_identity` rather than inventing one, so the fixture keeps describing the same dataset the rest of the snapshot describes. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> (cherry picked from commit e627dbb) * Document obsolete-format removals and cleanup validation * Record passing restacked plan, storage and serving checks * Remove generated evaluation archives and keep validation summaries * Normalize validation summary formatting * Remove PR process reports and defer execution design to shared-runtime PR * Keep discovery and calibration on dataset-bound snapshot version 3 * Use typed deployment configuration in process E2E fixtures * Remove unused streaming-config fixture after physical-only startup * Align process assertions with dataset-bound SDS and cold successor activation * fix: execute typed Planner fragments and preserve terminal resource errors * fix: bind exact integer samples without losing input types * test: retain resource failures across summary DAG execution * fix: treat cooperative query yields as pending execution * test: verify exact integer protocol input binding * refactor: name per-boundary physical fragments explicitly * fix: consume finalized Planner query candidate outputs * test: expose terminal error masking during revision changes * test: require Planner filters for bound protocol vectors * test: use Planner schema lifting module * fix: bind Planner filters and finalized query results * fix: preserve terminal execution errors before revision fencing * docs: describe bound Planner filter execution * refactor: return typed execution error directly * fix: retain state binding and sharing beneath explicit query readouts * fix: pin Planner query finalization for every candidate entry point * docs: distinguish query values from stored accumulator boundaries * test: assert exact Count state beneath its query readout * fix: preserve maintenance failures and stop failed worker admission * docs: specify bounded revisions and consistent query snapshots * feat: execute bounded precompute revisions with consistent SDS reads * fix: enforce query-wide resource and snapshot contracts * test: reject deployment candidates without snapshot agreement * docs: clarify local versus external query input candidates * test: retain Float64 formatting in external-only SQL oracle * docs: clarify precomputation terminology and remove stale review notes * docs: structure query DAG design around ownership and execution contracts * docs: align query execution layers with Planner physical candidates --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
Deployments need one coherent installed plan that binds queries and producers to stored state with known meaning. Routing fingerprints alone cannot distinguish equal metric expressions over different logical datasets, or represent two deployed outputs with the same semantics safely.
Before this PR
Backend configuration mixes plan concerns and does not establish the SDS identity contract from #737. A query binding cannot independently express the semantic computation and the authorized deployed output.
After this PR
Install a coherent
PrecomputePlan,QueryPlan, transmission configuration and immutable catalog snapshot. Planner exports the persisted output's canonical typed dependency closure, including logical dataset identity; deployment binding connects it to a stored output.Two
KLL(latency)outputs named hot/rebuild may share a definition but are independently bound. Another dataset changes the definition; relocating the same dataset preserves it. Same-version restart restores valid state. A new plan version remains cold until fresh input is produced and never adopts previous-version payloads.Catalog schema 6 and planning snapshot schema 3 establish this identity contract here. Typed Count/Rate support formerly in #771 is included because those semantics must remain distinct. Shared physical execution and its precompute execution design document continue in #774.
Obsolete installed-plan aliases, untyped projection decoders, the unused materialization wire adapter, and cross-version adoption metadata are removed. Recovery accepts only the current sidecar schema and propagates malformed or unsupported metadata errors. Producer partition rosters are explicit on the wire.
Dependency changes
asap_sketch_codeccontains the envelope encoding/decoding helpers used by backend ingest, DDSketch/KLL accumulators and wire-format tests. It supports removing theasap-precompute-rsCollector runtime dependency; it adds no sketch algorithms. The workspace member and data-plane dependency are intentional. With the Collector dependency removed, its Sketchlib patch block is unused and is removed too.cd7e9e0tobccc837because this implementation usesLogicalDatasetIdentityandSummarySemanticFragment::from_stored_output_in_dataset. Both APIs are absent at the old revision. All four Planner dependencies use the same immutable revision.Validation and scope
The SDS implementation passed locally: 117 type-library tests, 431 control-plane library tests, 1,165 data-plane library tests, 8 HTTP API tests, 13 serving integration tests, the production-process restart/warm-up test, strict all-target Clippy and formatting.
Coverage includes dataset changes, endpoint relocation, independent hot/rebuild bindings, tampered definitions, writer/read consistency, same-version restart without re-ingestion, and new-version warm-up. Earlier downstream validation also passed Level 1, exhaustive synthetic Level 2 selection, 332 storage tests and 13 serving integration tests.
Discovery, calibration and workload replay now preserve version-3 dataset identity. Process fixtures use typed deployment configuration instead of removed flat aggregation documents. At head
c11a540d, the complete local workspace run passes all 1,794 tests, including 18 compatibility process tests, four sketch oracle tests, the control-plane → data-plane process test, and cold-successor activation. The 81 Python tests, workspace formatting and strict all-target Clippy also pass. External-service opt-in tests were not enabled; this is not manual deployment verification.PR-specific evidence directories and generated archives are removed from the repository. Test summaries belong here; raw logs remain local or in CI.
Ad-hoc discovery, cross-version adoption and production-cost validation remain out of scope. Restricted native configuration helpers support explicit imported-state fixtures; production Planner compilation requires dataset identity. No manual deployment or human approval is claimed.